Papers by Rao Muhammad Anwer
Time Travel: A Comprehensive Benchmark to Evaluate LMMs on Historical and Cultural Artifacts (2025.findings-acl)
Copied to clipboard
Sara Ghaboura, Ketan Pravin More, Ritesh Thawkar, Wafa Al Ghallabi, Omkar Thawakar, Fahad Shahbaz Khan, Hisham Cholakkal, Salman Khan, Rao Muhammad Anwer
| Challenge: | TimeTravel is a benchmark of 10,250 expert-verified historical artifact samples spanning 266 distinct cultures across 10 major historical regions. |
| Approach: | They evaluate contemporary AI models on TimeTravel, highlighting their strengths and identifying areas for improvement. |
| Outcome: | The timeTravel benchmark covers 266 cultures and 10 major historical regions and aims to establish AI as reliable partner in preserving cultural heritage. |
LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs (2025.findings-acl)
Copied to clipboard
Omkar Thawakar, Dinura Dissanayake, Ketan Pravin More, Ritesh Thawkar, Ahmed Heakl, Noor Ahsan, Yuhao Li, Ilmuz Zaman Mohammed Zumri, Jean Lahoud, Rao Muhammad Anwer, Hisham Cholakkal, Ivan Laptev, Mubarak Shah, Fahad Shahbaz Khan, Salman Khan
| Challenge: | Existing approaches do not emphasize step-wise problem-solving. |
| Approach: | They propose a visual reasoning chain benchmark and a fine-grained reasoning metric that evaluates correctness and logical coherence at each step. |
| Outcome: | The proposed framework outperforms existing models in six benchmarks and is 5x faster during inference scaling. |
MAviS: A Multimodal Conversational Assistant For Avian Species (2025.emnlp-main)
Copied to clipboard
Yevheniia Kryklyvets, Mohammed Irfan Kurpath, Sahal Shaji Mullappilly, Jinxing Zhou, Fahad Shahbaz Khan, Rao Muhammad Anwer, Salman Khan, Hisham Cholakkal
| Challenge: | Existing multimodal large language models face challenges when it comes to specialized topics like avian species. |
| Approach: | They propose a large-scale multimodal avian species dataset that integrates image, audio, and text modalities for over 1,000 bird species. |
| Outcome: | The proposed model outperforms the baseline MiniCPM-o-2.6 by a large margin. |
BiMediX2 : Bio-Medical EXpert LMM for Diverse Medical Modalities (2025.findings-emnlp)
Copied to clipboard
Sahal Shaji Mullappilly, Mohammed Irfan Kurpath, Sara Pieri, Saeed Yahya Alseiari, Shanavas Cholakkal, Khaled M Aldahmani, Fahad Shahbaz Khan, Rao Muhammad Anwer, Salman Khan, Timothy Baldwin, Hisham Cholakkal
| Challenge: | BiMediX2 is a bilingual (Arabic-English) large multimodal model that supports text-based and image-based medical interactions. |
| Approach: | They introduce BiMediX2, a bilingual (Arabic-English) Bio-Medical EXpert Large Multimodal Model that supports text-based and image-based medical interactions. |
| Outcome: | The model outperforms existing models by over 9% in English and more than 20% in Arabic evaluations. |
LLMVoX: Autoregressive Streaming Text-to-Speech Model for Any LLM (2025.findings-acl)
Copied to clipboard
Sambal Shikhar, Mohammed Irfan Kurpath, Sahal Shaji Mullappilly, Jean Lahoud, Fahad Shahbaz Khan, Rao Muhammad Anwer, Salman Khan, Hisham Cholakkal
| Challenge: | Existing speech-enabled LLMs degrade conversational quality by modifying the LLM, compromising its linguistic capabilities. |
| Approach: | They propose a lightweight 30M-parameter, LLM-agnostic, autoregressive streaming TTS system that generates high-quality speech with low latency. |
| Outcome: | The proposed system achieves a significantly lower word error rate compared to speech-enabled LLMs while operating at comparable latency. |
CAMEL-Bench: A Comprehensive Arabic LMM Benchmark (2025.findings-naacl)
Copied to clipboard
Sara Ghaboura, Ahmed Heakl, Omkar Thawakar, Ali Husain Salem Abdulla Alharthi, Ines Riahi, Abduljalil Radman, Jorma Laaksonen, Fahad Shahbaz Khan, Salman Khan, Rao Muhammad Anwer
| Challenge: | Recent years have witnessed a significant interest in developing large multimodal models capable of performing various visual reasoning and understanding tasks. |
| Approach: | They propose to use Arabic as a language to evaluate large multi-modal models capable of performing visual reasoning and understanding tasks. |
| Outcome: | The proposed benchmark comprises eight diverse domains and 38 sub-domains to represent a large population of over 400 million speakers. |
Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMs (2025.emnlp-main)
Copied to clipboard
Wafa Al Ghallabi, Ritesh Thawkar, Sara Ghaboura, Ketan Pravin More, Omkar Thawakar, Hisham Cholakkal, Salman Khan, Rao Muhammad Anwer
| Challenge: | a benchmark is designed to assess the comprehension of Arabic poetry by large language models in 12 historical eras. |
| Approach: | They propose a benchmark to assess the comprehension of Arabic poetry by large language models in 12 historical eras. |
| Outcome: | The benchmark assesses the comprehension of Arabic poetry by large language models in 12 historical eras. |
DuwatBench: Bridging Language and Visual Heritage through an Arabic Calligraphy Benchmark for Multimodal Understanding (2026.eacl-long)
Copied to clipboard
Shubham Patle, Sara Ghaboura, Hania Tariq, Mohammad Usman Khan, Omkar Thawakar, Rao Muhammad Anwer, Salman Khan
| Challenge: | a benchmark of 1,272 samples containing about 1,475 unique words is available for Arabic calligraphy . the dataset reflects real-world challenges in Arabic writing, such as calligraphic variation and artistic distortions . |
| Approach: | They evaluated 13 leading Arabic and multilingual multimodal models and paired them with sentence-level annotations to evaluate their calligraphy models. |
| Outcome: | The benchmark evaluates 13 leading Arabic and multilingual multimodal models . it shows they struggle with calligraphic variation, artistic distortions, and precise visual–text alignment. |